Papers by Joon Son Chung

2 papers
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing speech translation approaches often overlook the transfer of speech patterns, leading to mismatches with source speech and limiting their suitability for dubbing applications.
Approach: They propose a diffusion-based speech-to-unit translation model with explicit duration control that enables time-aligned translation.
Outcome: The proposed system preserves key characteristics such as duration, speaker identity, and speaking speed while maintaining key characteristics.
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances have introduced GER frameworks that utilize LLMs to refine ASR outputs.
Approach: They propose a framework that allows a large language model to compose independent N-best hypotheses from separate automatic speech recognition (ASR) and visual speech recognition models.
Outcome: The proposed framework achieves 57.7% error rate gain over standard ASR baseline, compared to single-stream approaches that achieve only 10% gain.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations